Papers with inter-rater reliability

6 papers
HOPE: A Task-Oriented and Human-Centric Evaluation Framework Using Professional Post-Editing Towards More Effective MT Evaluation (2022.lrec-1)

Copied to clipboard

Challenge: Existing automated evaluation metrics for machine translation are expensive and lack inter-rater reliability.
Approach: They propose a task-oriented and human-centric evaluation framework for machine translation output based on professional post-e diting annotations.
Outcome: The proposed framework improves translation quality and system performance and transparency . it is cost-effective, easy to use and faster to implement .
Generation, Distillation and Evaluation of Motivational Interviewing-Style Reflections with a Foundational Language Model (2024.eacl-long)

Copied to clipboard

Challenge: Motivational Interviewing (MI) is a counselling technique used to guide people towards behaviour change.
Approach: They propose a method for distilling reflections from a foundational language model into smaller models that can be owned and controlled.
Outcome: The proposed method achieves 100% success rate on hold-out test set and 90% on the GPT-2 XL.
Development and Benchmarking of a Blended Human-AI Qualitative Research Assistant (2026.acl-industry)

Copied to clipboard

Challenge: Qualitative research emphasizes constructing meaning through iterative engagement with textual data.
Approach: They present and benchmark a qualitative research assistant system that allows researchers to identify themes and annotate datasets.
Outcome: The proposed system achieves an inter-rater reliability between Muse and humans of Cohen’s = 0.7 for well-specified codes.
Cross-replication Reliability - An Empirical Approach to Interpreting Inter-rater Reliability (2021.acl-long)

Copied to clipboard

Challenge: Respectable journals typically require reporting quantitative evidence for inter-rater reliability (IRR) of the data.
Approach: They propose to benchmark IRR against baseline measures in a replication dataset and use Cohen's (1960) kappa to measure inter-rater reliability.
Outcome: The proposed framework can be used to measure the quality of crowdsourced datasets.
French Tweet Corpus for Automatic Stance Detection (2020.lrec-1)

Copied to clipboard

Challenge: a new corpus of tweets is being developed for automatic stance detection of fake news . the task involves determining the attitude expressed in a text toward a target . this is a difficult task to overcome as discussions about fake news are controversial .
Approach: They propose to build a human-annotated corpus for automatic stance detection of tweets in french . they propose to use four classes broadly adopted by the community for annotation .
Outcome: The proposed corpus is the first freely available stance annotated tweet corpus in the french language.
Can Large Language Models Outperform Non-Experts in Poetry Evaluation? A Comparative Study Using the Consensual Assessment Technique (2025.emnlp-main)

Copied to clipboard

Challenge: Consensual Assessment Technique (CAT) for large language models is used to evaluate creativity, but is costly and time-consuming with non-experts.
Approach: They adapt the Consensual Assessment Technique (CAT) for Large Language Models to a 90-poem dataset with a ground truth based on publication venue.
Outcome: The proposed method outperforms the best human non-expert evaluations by significantly outperforming the best language models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations